Restoring Warped Document Image Based on Text Line Correction
نویسندگان
چکیده
Abstract Document images captured by camera often suffer from warping and distortions because of the bounded volumes and complex environment light source. These effects not only reduce the document readability but also the OCR recognition performance. In this paper, we propose a method to combine non-linear and linear compensation for correcting distortions of document images. First, due to the broken text result of Otsu binarization, an image preprocessing is used to remove the effect of background light. Second, the dewarping method using the cubic polynomial fitting equation is proposed to find out the optimal approximate text line for vertical direction rectification. Finally, we use linear compensation for horizontal direction rectification. Experimental results demonstrate the robustness of the proposed methodology and improve the accuracy rate of OCR recognition.
منابع مشابه
Document Image Dewarping Based on Text Line Detection and Surface Modeling (RESEARCH NOTE)
Document images produced by scanner or digital camera, usually suffer from geometric and photometric distortions. Both of them deteriorate the performance of OCR systems. In this paper, we present a novel method to compensate for undesirable geometric distortions aiming to improve OCR results. Our methodology is based on finding text lines by dynamic local connectivity map and then applying a l...
متن کاملRestoration of Arbitrarily Warped Document Images Based on Text Line and Word Detection
This paper presents a novel technique for efficient restoration of arbitrarily warped document images. Our aim is to recover document images that are mainly bounded volumes captured by a digital camera and suffer from non-linear warp. The proposed technique is applied on gray scale document images and is based on several distinct steps: an adaptive document image binarization, a text line and w...
متن کاملFast Restoration of Warped Document Image based on Text Rectangle Area Segmentation
The warp problems usually make the documents being hardly recognized. Specifically, when we copy a page of a thick book or bound document by digital photocopier, the resulted image is usually warped because of the thickness of the document. We focus on this problem and propose a fast method to restore the warped document image in this paper. The text rectangle area of the document is one of the...
متن کاملرفع اعوجاج هندسی متون بهکمک اطلاعات هندسی خطوط متن
Document images produced by scanners or digital cameras usually have photometric and geometric distortions. If either of these effects distorts document, recognition of words from such a document image using OCR is subject to errors. In this paper we propose a novel approach to significantly remove geometric distortion from document images. In this method first we extract document lines from do...
متن کاملDocument Analysis And Classification Based On Passing Window
In this paper we present Document analysis and classification system to segment and classify contents of Arabic document images. This system includes preprocessing, document segmentation, feature extraction and document classification. A document image is enhanced in the preprocessing by removing noise, binarization, and detecting and correcting image skew. In document segmentation, an algorith...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2014